Papers by Derry Tanti Wijaya

19 papers
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)

Copied to clipboard

Challenge: Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings.
Approach: They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research.
Outcome: The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings.
Did that happen? Predicting Social Media Posts that are Indicative of what happened in a scene: A case study of a TV show (2022.lrec-1)

Copied to clipboard

Challenge: Prior work identified and summarized scenes associated with a TV show by selecting a few representative social media posts (5 posts) that were published during the timeline of the scenes.
Approach: They propose a method to predict social media posts associated with a TV show from those that are not-indicative.
Outcome: The proposed method can predict posts indicative of what happened in a scene from those that are not-indicative based on high AUC's on social media posts associated with a popular TV show .
IndoCollex: A Testbed for Morphological Transformation of Indonesian Colloquial Words (2021.findings-acl)

Copied to clipboard

Challenge: Existing research on word normalization in Indonesian language relies on static dictionaries and machine translation.
Approach: They propose to use Twitter to annotate Indonesian colloquial words with their standard forms and their word formation types/tags to perform morphological word normalization.
Outcome: The proposed dataset analyzes morphological word normalization on Indonesian colloquial Lexicons and provides a baseline for future work.
Cultural and Geographical Influences on Image Translatability of Words across Languages (2021.naacl-main)

Copied to clipboard

Challenge: Neural machine translation models produce poor translations when there are few/no parallel sentences to train the models.
Approach: They define image translatability as the translability of words as images associated with words in different languages that have a high degree of visual similarity.
Outcome: The proposed model improves upon text-only models only marginally.
BU-NEmo: an Affective Dataset of Gun Violence News (2022.lrec-1)

Copied to clipboard

Challenge: Using a dataset that contains headline and image pairings from 840 news articles, we explore the relationship between image and text influence on human emotional response.
Approach: They propose to use a U.S. gun violence news dataset that contains headline and image pairings from 840 news articles with 15K high-quality crowdsourced annotations on emotional responses.
Outcome: The proposed dataset includes annotations on the dominant emotion experienced with the content, the intensity of the selected emotion and an open-ended, written component.
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information (2025.findings-acl)

Copied to clipboard

Challenge: Prior research has focused on toxicity and polarization as separate problems . extreme polarizing deepens divisions, often leading to hostility and fragmentation .
Approach: They propose to use a multi-label Indonesian dataset annotated for toxicity, polarization, and annotator demographic information to study polarizing language and toxicity.
Outcome: The proposed dataset shows that polarization cues improve toxicity classification and vice versa.
“Wikily” Supervised Neural Translation Tailored to Cross-Lingual Tasks (2021.emnlp-main)

Copied to clipboard

Challenge: Unsupervised neural machine translation models perform well in low-resource or distant languages.
Approach: They propose a model that leverages Wikipedia for machine translation and cross-lingual tasks without supervision from external parallel data or supervised models in target language.
Outcome: The proposed model outperforms supervised models in Arabic and English translation tasks.
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines (2025.naacl-long)

Copied to clipboard

Challenge: Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts.
Approach: They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset.
Outcome: The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages.
Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for classifying discrete emotions from news headlines have been limited to using headlines.
Approach: They propose to use people’s free-text explanations to classify emotions elicited by news headlines to generate emotion explanations from headlines.
Outcome: The proposed method improves on methods that only use headlines and train a pretrained model for explanation generation.
Do Language Models Understand Honorific Systems in Javanese? (2025.acl-long)

Copied to clipboard

Challenge: Despite its cultural and linguistic significance, there has been limited progress in developing a comprehensive corpus to capture these variations for natural language processing (NLP) tasks.
Approach: They propose to use a dataset to capture the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework, to assess the ability of language models to process various levels of Javanesi honorifics.
Outcome: The proposed dataset encapsulates the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework.
What Do Indonesians Really Need from Language Technology? A Nationwide Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Despite efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers.
Approach: They conduct a nationwide survey to assess the actual needs of native Indonesian speakers.
Outcome: The findings indicate that addressing language barriers is the most critical priority . concerns around privacy, bias, and the use of public data highlight the need for greater transparency and clear communication to support broader AI adoption.
OpenFraming: Open-sourced Tool for Computational Framing Analysis of Multilingual Data (2021.emnlp-demo)

Copied to clipboard

Challenge: Existing frameworks for analyzing frames in multilingual text documents are available online and via an API.
Approach: They propose a web-based system for analyzing frames in multilingual text documents . framework combines unsupervised and supervised machine learning and leverages a state-of-the-art multilingual language model .
Outcome: The proposed framework can significantly improve frame prediction performance while requiring a small sample of manual annotations.
Mitigating Translationese in Low-resource Languages: The Storyboard Approach (2024.lrec-main)

Copied to clipboard

Challenge: Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which introduce the translationese effect.
Approach: They propose a method that uses storyboards to elicit more fluent and natural sentences from native speakers without direct exposure to the source text.
Outcome: The proposed method compared with traditional translation-based methods in terms of accuracy and fluency.
Multi-Label and Multilingual News Framing Analysis (2020.acl-main)

Copied to clipboard

Challenge: Recent studies have focused on news framing in English, but few studies have explored how it can be extended to other languages and in multi-label settings.
Approach: They propose a method that leverages dictionary and few annotations to detect frames from just the headline in a low-resource context.
Outcome: The proposed method performs better than translating the entire headline to the source language . it can be scaled up to many languages, even those without existing translation technologies .
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts (2025.acl-long)

Copied to clipboard

Challenge: NusaAksara covers 8 scripts across 7 languages, including low-resource languages not commonly seen in NLP benchmarks.
Approach: They propose a benchmark for Indonesian scripts that includes their original scripts and a dataset that includes 8 scripts across 7 languages.
Outcome: The proposed benchmark covers 8 scripts across 7 languages, including low-resource languages not commonly seen in NLP benchmarks.
AugCSE: Contrastive Sentence Embedding with Diverse Augmentations (2022.aacl-main)

Copied to clipboard

Challenge: Similar work has shown that a single augmentation can be used to learn a robust generalpurpose representation with contrastive learning.
Approach: They propose a unified framework to utilize diverse sets of data augmentations to achieve a better, general-purpose sentence embedding model.
Outcome: The proposed framework achieves state-of-the-art results on downstream transfer tasks and performs competitively on semantic textual similarity tasks, using only unsupervised data.
Prediction of People’s Emotional Response towards Multi-modal News (2022.aacl-main)

Copied to clipboard

Challenge: BU-NEmo dataset extends from 320 to 1,297 news headline and lead image pairings and collects 38,910 annotations in a crowdsourcing experiment.
Approach: They extend the U.S. gun violence news-to-emotions dataset from 320 to 1,297 news headline and lead image pairings and collect annotations in a crowdsourcing experiment.
Outcome: The proposed models outperform baseline models on the NEmo+ dataset by large margins across several metrics.
Detecting Frames in News Headlines and Lead Images in U.S. Gun Violence Coverage (2021.findings-emnlp)

Copied to clipboard

Challenge: Journalists have been using both text and images to frame news stories . lead images may carry additional background knowledge about the event .
Approach: They find that combining lead images and contextual information with text improves news framing . they release the first multimodal news framming dataset related to gun violence in the u.s.
Outcome: The study shows that combining lead images with text improves prediction of news frames . it also shows that using multiple modes of information improves frame image relevance .
RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs (2023.acl-long)

Copied to clipboard

Challenge: Despite their success, even the largest language models make mistakes.
Approach: They propose a framework where one language model can generate critiques to improve its peer's performance.
Outcome: The proposed framework improves the performance of a fixed model 200 times its size by 10% over other models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations